JAMA Pediatrics
● American Medical Association (AMA)
Preprints posted in the last 7 days, ranked by how well they match JAMA Pediatrics's content profile, based on 10 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.
Show abstract
Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.
Ebneabbasi, A.; Warrier, V.; Montagnese, M.; Romero Garcia, R.; Bethlehem, R. A. I.; Rittman, T.
Show abstract
Neighbourhood deprivation is one of the few potential policy-modifiable risk factors for psychiatric and neurological disorders, but the neurobiological pathways underlying these associations remain unclear. We investigated these relationships across three cohorts spanning the life span: the Healthy Brain and Child Development (HBCD) Study (n = 84, aged 0 to 4 weeks postnatal), the Adolescent Brain Cognitive Development (ABCD) Study (n = 4,792, aged 9 to 10 years), and the UK Biobank (UKB; approximately 500,000 adults, aged 44 to 87 years). Neighbourhood deprivation was associated with elevated disease risk, and individual lifestyle factors accounted for only a small fraction of this burden, indicating that the much larger residual effect reflects broader contextual characteristics of deprived environments rather than individual behaviours alone. Across all cohorts, greater deprivation consistently predicted lower cortical and subcortical brain volume, with effects detectable in early development and substantially stronger in adulthood. Across disorders, regional brain volume emerged as a consistent neuroanatomical mediator linking neighbourhood deprivation to neuropsychiatric disease. We further showed that deprivation preferentially affects brain regions intrinsically vulnerable to neuropsychiatric disorders. Spatial decoding analyses implicated dopaminergic and serotonergic neurotransmitter systems together with specific excitatory and inhibitory neuronal classes. Importantly, both the deprivation effects and their neuroanatomical mediation patterns were replicated across independent populations. Our study delivers a translational framework linking neighbourhood deprivation to brain health, which could inform public health policies and preventive interventions.
Chen, Y.; Puckett, H.; Clarot, G.; Hawkins, B.; Sharp, K.; Todd, D. A.; Lopez, A.; Bertollo, J. R.; Behar, H. E.; Zeithamova, D.; Xie, H.; Verbalis, A.; VanMeter, A. S.; Gaillard, W. D.; Kenworthy, L.; Vaidya, C. J.
Show abstract
Generalization is a key cognitive process that allows humans to flexibly apply prior knowledge to guide new behaviors. Difficulties with generalization and flexibility are observed across neurodevelopmental disorders, especially autism, limiting adaptive function and quality of life. Cognitive-behavioral treatment benefits some but not all autistic individuals. As treatment requires application of learned skills to everyday life, variability in generalization ability may limit intervention success in autism. While cognitive substrates of learning and generalization are well established, their potential for explaining clinical outcomes is not known. Here, we combined a category learning task with computational modelling to distinguish two learning strategies underlying generalization -- prototype abstraction vs. exemplar memorization -- and tested whether individual differences in these learning strategies predicted real-world intervention outcomes in autistic youth. Fifty-four participants completed the category learning task at two pre-intervention timepoints, and then completed Unstuck and On Target:14-22 intervention targeting flexible problem solving, goal setting, and planning. We found that participants who consistently relied on prototype abstraction (N=26) were subsequently more likely to benefit from the intervention, showing improvement in parent- and self-reported flexibility. These findings identify prototype abstraction as a clinically relevant cognitive capacity that may help explain individual differences in intervention response and support the tailoring of interventions. More broadly, they demonstrate the value of linking basic cognitive mechanisms to clinical outcomes and may inform strategies to enhance the effectiveness of cognitive-behavioral interventions for youth with developmental disabilities.
Humphries, C.; Brett, J.; Gruber, F.; James, E.; McKendrick, T. I.; McNairn, K. C.; Miell, A.; O'Brien, R.; Rahman, F.; Schölin, L.; Stewart, M.; Casey, A.
Show abstract
Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients [≥]16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Leuenberger, L. M.; Shoman, Y.; Romero, F.; Sasaki, M.; Deligianni, X.; Goebel, N.; Mozun, R.; Bielicki, J. A.; Burckhardt, M.-A.; Saner, C.; Schwitzgebel, V.; Hauschild, M.; Righini Grunder, F.; Mueller, P.; Schlapbach, L. J.; Jenni, O.; Spycher, B. D.; Kuehni, C. E.; Belle, F. N.; SwissPedHealth consotrium,
Show abstract
BACKGROUND: We used anthropometric data from electronic health records (EHRs) of Swiss childrens hospitals to evaluate growth references and estimate centile curves. METHODS: We received EHRs extracted from seven Swiss childrens hospitals and analysed two samples: all children with a height, weight, body mass index (BMI), or head circumference recording, and a subsample restricted to children without diseases potentially affecting growth, weighted to represent the general population. We calculated mean z-scores based on the World Health Organization growth references adopted for Switzerland in 2011 (CH-WHO 2011) and current Swiss growth references (Swiss 2026). We estimated sex-specific centile curves in the subsample using generalised additive models for location, scale, and shape. RESULTS: We included 213,868 children with height, 448,002 with weight, 209,244 with BMI, and 67,397 with head circumference recordings. Mean z-scores in the all children sample were (CH-WHO 2011; Swiss 2026): height (0.10; -0.19), weight (0.16; -0.09), BMI (0.04; -0.07), head circumference (-0.28, -0.28); and in the subsample: height (0.34; 0.00), weight (0.27; 0.01), BMI (0.18; 0.05), and head circumference (0.04; 0.01). The 50th height, weight, BMI, and head circumference centiles of girls and boys in the subsample closely followed those of Swiss 2026, with slightly wider 3rd and 97th centiles in infancy and adolescence. CONCLUSION: Height, weight, BMI, and head circumference centiles aligned well with the Swiss 2026 growth references in Switzerland, demonstrating that hospital EHRs could contribute to future growth references.
Quigley, H.; Gardiner, B.; McDaid, L.; O'Donnell, C.
Show abstract
Autism Spectrum Disorder (ASD) is a heterogeneous neurodevelopmental condition defined by differences in social communication and restricted, repetitive behaviours. As diagnostic criteria have broadened, ASD is now recognised across a wider range of individuals, raising key questions about its structure: does ASD have discrete sub-types, or is it better conceptualised as a continuous, possibly multidimensional, condition? We aim to explore whether a multidimensional continuum model more accurately captures the variability within ASD. We analysed a large SPARK phenotypic dataset of medical history and diagnostic surveys (background history, SCQ, RBS-R; n=36,710 individuals). We apply and compare two traditional statistical approaches, Factor Analysis and Gaussian Mixture Models, with a modern machine learning technique, the Variational Autoencoder (VAE). VAEs reconstructed unseen test data with ~4-fold better accuracy than Factor Analysis, and ~8-fold better accuracy than Gaussian Mixture Models. We identified four stable latent factors across 100 independently trained VAEs. These four dimensions provide an individual behavioural profile that can be visualized using radar-plots, offering a compact way to compare profiles at the person level. Through further analysis, we found evidence for 3 overlapping clusters or subtypes of ASD identified within the 4D latent space. This work aims to inform new ways of modelling ASD using a VAE that will be able to discern between a continuum or a clustered output and that go beyond binary diagnosis, instead reflecting the complex range of trait profiles, with implications for personalised diagnosis and intervention.
Li, D.; Feng, Q.; Zhang, Y.; Chen, H.; Wang, X.; Shen, C.
Show abstract
Background National childhood respiratory pathogen spectra are diversifying nearly everywhere - within-country diversity rose in 203 of 204 countries between 1990 and 2023 - yet whether countries are diversifying toward a common spectrum or along divergent paths is unknown. We quantified between-country compositional distance of national pathogen spectra over the same period. Methods We built national pathogen share vectors from Global Burden of Disease Study 2023 lower respiratory infection etiologic attributions (26 pathogens, 204 countries, ages 0-19 years) at five timepoints spanning 1990-2023. Between-country distance was measured as all pairwise Jensen-Shannon divergences (JSD; primary) and Bray-Curtis dissimilarities, with Baselga and Jaccard decompositions; robustness was assessed across metrics, pathogen panels, low-count thresholds and a balanced panel of 107 countries. Results Mean pairwise JSD rose from 0.0084 in 1990 to 0.0283 in 2023 (+238%; trend p = 0.030), peaking in 2021 (+283%) with a partial 2023 pullback. Bray-Curtis dissimilarity rose +120% and the balanced panel +423%. Divergence was entirely balanced variation (share reallocation), with spectrum richness rising from 18.5 to 21.1 of 26 pathogens. Dispersion rose fastest for influenza (coefficient of variation 0.03 to 0.55) and respiratory syncytial virus (0.08 to 0.48). Within-region distance rose in every computable GBD super-region (five of seven): divergence occurs within regions, not between blocs. Conclusions National spectra are re-sorting along country-specific axes as vaccine-preventable dominance recedes at different speeds. Diversification is universal, but convergence is absent: the transition at the etiologic-spectrum level is asynchronous and path-dependent, with implications for empirical treatment policy and pathogen surveillance.
Li, D.; Chen, H.; Xie, J.; Li, J.; Wang, X.; Shen, C.
Show abstract
Background The historic decline in childhood pneumonia mortality was driven substantially by single-pathogen vaccines against Haemophilus influenzae type b (Hib) and Streptococcus pneumoniae. Yet the pathogen spectrum underlying child pneumonia deaths is diversifying: the effective number of pathogens rose from 5.57 in 1990 to 9.94 in 2023, and the residual burden is shifting toward opportunistic and hospital-associated pathogens for which no licensed childhood vaccines exist. This paper asks how resources should be sequenced between single-pathogen interventions and platform investments as this transition proceeds. Methods We analyzed Global Burden of Disease Study 2023 deaths from 29 pathogens in ages 0-19 years by super-region, combined with WHO/UNICEF Estimates of National Immunization Coverage (WUENIC) for PCV3 and Hib3. We quantified the spectrum transition under two denominators (26- and 29-pathogen calibers), constructed a share-by-intervenability matrix assigning each pathogen to a dominant intervention channel (vaccine-reachable, mixed, platform-sensitive) under explicit classification rules, compared platform-sensitive deaths with a transparently computed scenario of residual vaccine-preventable deaths, and cross-classified pathogens by age tropism and poverty lock. We anchored platform interventions to verified published evidence. Results The vaccine-preventable group share fell from 54.0% to 40.2% while the opportunistic/hospital group rose from 18.1% to 23.1% (29-pathogen caliber, 1990-2023). Super-region vaccine coverage showed no significant association with pathogen-share change (PCV3 Spearman rho = 0.108, p = 0.818; Hib3 rho = -0.036, p = 0.939), a null result we report as evidence that simple coverage-burden correlations do not hold at the regional level, not as evidence against vaccine value. In 2023, vaccine-reachable pathogens accounted for 441,410 deaths (45.7%, channel including COVID-19), mixed for 126,926 (13.1%), and platform-sensitive pathogens for 396,995 (41.1%). Platform-sensitive deaths were 2.9-5.1 times the scenario estimate of residual vaccine-preventable deaths (52,435-77,512). Nine of 14 classifiable pathogens fell into the poverty-locked, infant-tropic cell (480,922 deaths; Fisher OR = 9.0, p = 0.1758). Conclusions The marginal value of single-pathogen strategies declines as the spectrum diversifies and residual deaths concentrate in platform-sensitive, poverty-locked, infant-tropic pathogens. Vaccine scale-up remains a certain and sizeable opportunity; the next increment of marginal resources should increasingly fund platform capabilities (oxygen systems, antimicrobial access and stewardship, infection prevention and control, referral, and nutrition) delivered as a package to the populations where the residual burden is locked.
Witham, M.; Evison, F.; Bellass, S.; Cooper, R.; Gallier, S.; Pretorius, S.; Sapey, E.; Suklan, J.; Sayer, A. A.
Show abstract
Study Objective Little is known about where in hospital care for multiple long-term conditions (MLTC) is delivered. We aimed to describe pathways of care (ward transfers) and outcomes for people admitted to hospital for unscheduled care by MLTC status and other key sociodemographic characteristics. Design and setting Analysis of routinely-collected electronic health records from a large acute UK hospital. Participants Adult unscheduled care admissions from 1st July 2018 to 30th June 2019. The presence of two or more of 59 long-term conditions was ascertained using ICD-10 codes from previous hospital discharges. Main outcome measures Markov state transition probabilities were derived for ward moves and compared for MLTC vs no MLTC, age, sex, ethnicity and neighbourhood deprivation. Outcomes (length of stay, death, readmission, move from definitive ward) and time spent in emergency and assessment departments were compared between subgroups. Results A total of 33,252 adults, mean age 56.0 (SD 21.9) years were analysed; 14,834 (42.4%) had MLTC. People with MLTC were more likely to die in hospital (4.2 vs 1.9%, p<0.001), transfer to internal medicine wards or older peoples medicine wards, were less likely to transfer to surgical wards, had longer median length of stay (1.83 vs 0.69 days, p<0.001), stayed longer in acute medical units (15.5 vs 9.6 hours, p<0.001), and were more likely to move from their definitive ward (18.2 vs 16.4%, p=0.002). Conclusion Unscheduled hospital care pathways are complex and differ for people with MLTC, who have worse outcomes and may be less likely to receive optimal care.
Liu, H.; Mizani, M. A.; Zhao, Y.; Wood, A.; Inouye, M.; Price, A. L.; Jiang, X.; CVD-COVID-UK/COVID-IMPACT Consortium,
Show abstract
Predicting disease risk from prior diagnoses is fundamental to clinical decision-making, particularly during health emergencies such as the COVID-19 pandemic, when individuals with long-term conditions may be disproportionately vulnerable to adverse outcomes. Despite intense interest in developing models to predict disease risk from prior diagnoses (1-3), most prediction models do not estimate effects of each prior diagnosis on disease risk conditional on other diagnoses, limiting interpretability and clinical utility. We developed the Comorbidity Risk Score (CRS), trained on 13 million individuals (age 40-69) from linked electronic health record (EHR) datasets of the entire population of England, to predict COVID-19 hospitalisation and 87 other disease outcomes. CRS was trained at close to saturated sample size and precisely estimated the effects of 212 prior diagnoses on the 88 disease outcomes, conditional on all other prior diagnoses. Correlations of CRS effect sizes across outcomes (e.g. 0.76 for myocardial infarction vs. hyperlipidaemia) matched the corresponding genetic correlations (e.g. 0.79 for myocardial infarction vs. hyperlipidaemia), confirming that comorbidity architectures capture disease aetiology. On average, CRS identified 5% of the population with 3.4-fold higher disease risk, including myocardial infarction (4.4-fold), lung cancer (6.5-fold), and COVID-19 hospitalisation (6.3-fold). Using prior diagnoses alone, CRS outperformed state-of-the-art clinical COVID-19 models (4). Furthermore, CRS (N=13 million) substantially outperformed state-of-the-art AI (1) (N=0.5 million) and linear (3) (N=0.5 million) models in predicting disease risk, suggesting that training sample size outweighs model complexity. CRS attained near-perfect transferability across self-reported ethnicities (e.g., Black vs. White: AUROC ratio = 97.3%). Finally, CRS distinguished independently predictive comorbidities from indirect associations, e.g., lipid metabolism disorder was a strong predictor of myocardial infarction risk but not ischaemic stroke, after conditioning on other prior diagnoses. In conclusion, CRS provides a comprehensive resource for understanding the impact of comorbidities on COVID-19 and other future diseases, revealing disease aetiology while enabling powerful prediction of disease risk.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Lee, H.-W.; Huang, Y.-H.; McAndrew, T. C.
Show abstract
Introduction. By the end of 2023, many low-income countries had not reached 50% COVID-19 vaccine coverage, while most high-income countries had exceeded 80%. It remains unclear whether receiving vaccine deliveries translated into faster population coverage. We examined cross-national inequalities in the timing of the vaccine rollout and whether deliveries through the COVID-19 Vaccines Global Access (COVAX) facility were associated with subsequent national uptake. Methods. We conducted an observational study of 218 countries and territories using country-level data up to December 2023. We used generalized additive mixed models to identify country-level correlates of coverage at an early and a later stage of the pandemic, survival analysis to compare the time to 50% coverage between COVAX Advance Market Commitment (AMC) and non-AMC countries, and an event study to estimate the association between the timing of the first COVAX delivery and subsequent monthly coverage in AMC countries. Results. AMC-supported countries reached 50% coverage substantially more slowly than non-AMC countries. The hazard of reaching the threshold was 0.17 times that of non-AMC countries at month 1 (95% CI 0.07 to 0.41) and 0.53 times at month 18 (95% CI 0.33 to 0.85). One year after rollout began, 65.9% of AMC countries (95% CI 56.7 to 76.6) had not reached 50% coverage, compared with 21.1% of non-AMC countries (95% CI 15.1 to 29.5). The timing of COVAX deliveries was not significantly associated with subsequent national uptake in any post-delivery month. In the early stage of rollout, higher maternal mortality was associated with lower coverage, while a larger urban population was associated with higher coverage. By the end of the observation period, larger household size was associated with lower coverage, while higher health expenditure and a larger urban population were associated with higher coverage. Conclusion. Receiving COVAX deliveries was not, on its own, associated with faster coverage. Coverage differences were more consistently associated with country-level structural and health-system characteristics, while we found no significant association with the timing of the first COVAX delivery. Achieving vaccine equality likely requires strengthening the capacity of health systems to convert deliveries into administered doses, and preparedness efforts should invest in last-mile delivery capacity ahead of future emergencies.
Bastien, J.; Garcia, K.; Wallace, A. L.; Sullivan, R. M.; Hoh, E.; Wade, N. E.
Show abstract
Background: As cannabis policy changes in the United States, secondhand cannabis smoke (SCS) is increasingly common, including within families. However, prevalence of exposure and clinical correlates over time in adolescents are not fully understood. Objectives: (1) To estimate the prevalence of SCS and personal cannabis use in US-based teens exposed to SCS, and (2) examine the cognitive trajectories of adolescents exposed to SCS compared to non-exposed peers. Methods: Data from the Adolescent Brain Cognitive Development (ABCD) Study was used. Participants (n=11,316 of full cohort with follow-up data; n=776 with self-reported family SCS exposure) attended yearly visits from ages 11-17, completing substance use interviews, toxicological testing, and the NIH Toolbox Cognitive battery. Youth with SCS but no personal cannabis use (n=419; 47% female) were matched on prenatal substance exposure, family substance use history, and sociodemographics to non-SCS exposed and non-cannabis-using youth with a 1:2 ratio (Controls n=838). Linear mixed-effects models assessed cognitive performance by SCS*age interactions, accounting for random effects of subject and family. Covariates included sex and alcohol, nicotine, and other substance use. Secondary models analyzed performance by cumulative waves of reported SCS exposure interacting with age. Results: Of the full cohort, 6.9% (n=776) reported exposure to SCS. Of these individuals, 46% endorsed lifetime personal cannabis use by age 17, relative to 20% of non-SCS exposed youth (OR=3.83[95%CI:3.29,4.44]). Within matched participants, SCS*age demonstrated a significant interaction on attention and inhibitory control ({beta}=-0.32, p=.028), with SCS demonstrating reduced improvement over time. More waves of exposure were also associated with worse performance over time ({beta}=-0.39, p=.057). Discussion: Almost half of those who had been exposed to SCS endorsed personal cannabis use. Cognitive findings were domain specific, similar to findings in secondhand tobacco: SCS exposed youth showed restricted improvement in attention and inhibitory control by age 17. Public health and policymakers should make efforts to curb youth SCS exposure, given the potential for risk which has not been fully explored to date.
Sörnyei, D.; Kovacs, F. M.; Benedek, T.; Ori, D.; Farkas, K.
Show abstract
The Autism Spectrum Quotient (AQ-50) is widely used to assess autistic traits, yet its Hungarian version has not been psychometrically evaluated. We assessed the reliability, factor structure, temporal stability, convergent validity, and clinical utility of the Hungarian AQ-50 and a revised translation (AQ-50-HU-R) in two samples (N1 = 1967; N2 = 423), including autistic and non-autistic participants. The AQ-50-HU-R showed high internal consistency and test-retest reliability. A bifactor model provided the best fit ({chi}2[1125] = 1650.433, p < 0.001; CFI = 0.991; TLI = 0.990; RMSEA = 0.033 [90% CI = 0.030-0.037]; SRMR = 0.083), with 71% of common variance attributable to a general autistic traits factor. The total score distinguished clinically verified autistic participants from participants reporting no ASD diagnosis (AUC = 0.906), with a cutoff of 25. Associations with ADOS scores were weak or nonsignificant. The AQ-50-HU-R is best interpreted as a reliable total-score screening measure, supporting referral for comprehensive autism assessment.
Li, D.; Xie, J.; Xue, J.; Chen, H.; Wang, X.; Shen, C.
Show abstract
Background Respiratory infections remain the leading infectious cause of death among children and adolescents, yet the share of these deaths that could be averted with currently feasible care is not routinely quantified. Existing amenable-mortality frameworks rely on cause lists and population-level mortality benchmarks and do not exploit information on how many episodes occur. We propose an episode-fatality-ratio (EFR) frontier approach and apply it to lower respiratory infections (LRI), whooping cough (pertussis) and upper respiratory infections (URI) in 204 countries, 1990-2023. Methods For each cause, country and year we computed EFR = deaths/incident episodes using Global Burden of Disease (GBD) 2023 estimates for ages 0-19 years. The frontier was defined as the 10th-percentile country EFR within each GBD super-region, cause and year; avoidable deaths = max(0, deaths - episodes x frontier EFR). Primary estimates are deterministic; 95% uncertainty intervals (UIs) come from 2,000 Monte Carlo draws. Sensitivity analyses varied the frontier percentile, applied an aspirational global frontier, constructed pertussis counterfactuals, and recomputed all estimates within the single under-5 age band. Results In 2023, 333,803 childhood deaths from lower respiratory infections (95% UI 289,123-417,460; 46.9% of LRI deaths) were avoidable. Summing the three causes deterministically gives 391,034 avoidable deaths (46.5% of 840,444); the combined figure is a deterministic sum, and a UI is available for the LRI component only. The pertussis (43,958; 39.0%) and URI (13,273; 81.0%) estimates are secondary: their deterministic point values fall below their own Monte Carlo intervals and the underlying death estimates carry very wide uncertainty (global pertussis UI 12,545-321,874). Avoidable deaths fell from 1,050,468 (44.9%) in 1990, but between 2019 and 2023 the avoidable share for LRI+URI barely moved (48.7% to 47.7%) while absolute avoidable deaths fell 14.5%, a pattern consistent with stalled convergence to the frontier. Sub-Saharan Africa plus South Asia held 73.1% of avoidable deaths in 2023 versus 41.8% in 1990; ten countries accounted for 59.1%. Conclusion Nearly half of childhood respiratory-infection deaths remain avoidable relative to within-region best practice, and the residual burden is increasingly concentrated in low-income settings. In the pertussis counterfactual, most countries kept pace with their regional frontier, so further gains require advancing the frontier itself through quality-of-care improvements.
Packard, S. E.; Russo, T.; Parrott, J.; Sisti, J.; Lans, A.
Show abstract
Objectives: To estimate the prevalence of Post-Exertional Malaise (PEM) among adults with prior COVID-19 and associated mental health and disability outcomes. Methods: We conducted a cross-sectional analysis of data from a survey of 9,620 adults with prior COVID-19 in New York City, collected May - June 2024. PEM was measured with the DePaul Symptom Questionnaire - Post Exertional Malaise, categorized by symptom duration (< 14 vs. [≥]14 hours). Weighted prevalence estimates were stratified by socio-demographic and clinical characteristics. Modified Poisson regression was used to assess the association of PEM with depression, anxiety, and disability. Results: The prevalence of PEM symptoms was 20.9% overall and 4.0% with symptom duration [≥]14 hours, representing over 800,000 New Yorkers affected and over 150,000 who meet a diagnostic criterion for ME/CFS. PEM prevalence was higher among women, transgender and non-binary adults, people of color, and lower educational attainment, chronic comorbidities, or disabilities. PEM was associated with 3 - 4 times higher prevalence of mental health outcomes and 4 - 5 times higher disability scores. Conclusions: PEM symptoms were common and strongly associated with disability and adverse mental health. Screening, pathways to care, and supportive policies are needed to mitigate long-term consequences, particularly among marginalized populations.
Chowdhury, A. R.; Chowdhury, B.
Show abstract
Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.
Manikam, L.; Fatima, A.; Patil, P.; Mayadewi, C. A.; El Khatib, T.; Drazdzewska, J.; Oyebode, O.; Llewellyn, C. H.; Webb-Martin, K.; Irish, C.; Archibong, M.; Gilmour, J.; Kalungi, P.; Batura, N.; Shringarpure, K.; Lakhanpaul, M.; Heys, M.; NEON Steering Team,
Show abstract
South Asian communities in the UK experience disproportionate maternal and child health inequalities linked to non-recommended infant feeding practices, limited health literacy, and socioeconomic constraints. Participatory learning and action (PLA) is effective in low- and middle-income countries, but high-income evidence is scarce. This pilot assessed the feasibility of a community facilitator-led PLA intervention to improve infant feeding among South Asian families in East London. A three-arm pilot feasibility cluster randomised controlled trial (ISRCTN10234623) was conducted in Tower Hamlets and Newham, East London (May-September 2022), with 12 wards randomised 1:1:1 to face-to-face PLA, online PLA, or usual care. Multilingual community facilitators delivered eight biweekly sessions over 14 weeks. Feasibility outcomes were assessed against prespecified Go/Stop criteria; exploratory outcomes included child feeding behaviours (Children's Eating Behaviour Questionnaire, CEBQ), parental feeding style (Parental Feeding Style Questionnaire, PFSQ), and child BMI Z-scores. Of 263 enrolled participants, 261 had a recorded trial arm allocation; consent to the pilot feasibility study was 70.7% (186/263; 95% CI 65.0-75.9%) meeting the [≥]50% Go criterion. Attendance was 37% (Tower Hamlets 59%, Newham 29%), below the [≥]80% Go threshold. Six-month retention was 54.8% (Tower Hamlets 78%, Newham 48.5%; 95% CI 41.8-55.3%), triggering the Definite Stop criterion. Significant baseline imbalances included BMI Z-score (p = 0.005), ethnicity, borough, and education; no between-arm BMI differences were observed at follow-up (p = 0.249). CEBQ and PFSQ baseline completion was 24.5% and 23.0%, with no usable follow-up data. PLA Phases 3 and 4 were not completed by any group; all participants providing feedback reported it acceptable. Recruitment was feasible and the intervention acceptable, but a Definite Stop criterion was triggered in Newham, no group completed the full PLA cycle, and outcome data were insufficient for evaluation. A definitive trial requires stratified randomisation, digitised multilingual data collection, participant reimbursement, and explicit PLA phase-completion criteria.
Nkulikwa, Z. A.
Show abstract
The analysis uses a global 2010-2023 panel comprising 3,038 economy-years across 217 economies. It explicitly separates between-economy and within-economy estimands and tests the longitudinal interpretation using an identical-sample temporal analysis with cluster-aware coefficient contrasts, a formal isometric log-ratio sensitivity analysis, independent fixed-effects replication, and wild-cluster-bootstrap inference. The central finding is deliberately calibrated: cross-economy agreement cannot validate national sugar availability for longitudinal obesity surveillance. The study identifies temporal and construct instability without claiming that sugar is protective or that the mechanisms producing the instability have been identified. The manuscript aligns well with PLOS ONEs emphasis on technically sound, transparent and reproducible research of broad relevance. All data required to reproduce the findings, complete metadata, executable code, full-precision results, diagnostic outputs and a completed STROBE checklist are provided as S1-S5. Figures are provided separately as compliant 350-dpi TIFF files. The study used only publicly available, aggregated economy-year statistics and involved no individual participants, identifiable information or biological specimens; institutional ethics review and consent were therefore not required. This is original work; it is not under consideration elsewhere, and the sole author has approved the submission and accepts responsibility for its content. Funding and competing-interest declarations will be entered accurately in the submission portal. An Academic Editor with expertise in nutritional epidemiology, global health metrics, longitudinal panel methods, or food-system surveillance would be well placed to assess the work.